shape and appearance
CtrlNeRF: The Generative Neural Radiation Fields for the Controllable Synthesis of High-fidelity 3D-Aware Images
The neural radiance field (NERF) advocates learning the continuous representation of 3D geometry through a multilayer perceptron (MLP). By integrating this into a generative model, the generative neural radiance field (GRAF) is capable of producing images from random noise z without 3D supervision. In practice, the shape and appearance are modeled by z_s and z_a, respectively, to manipulate them separately during inference. However, it is challenging to represent multiple scenes using a solitary MLP and precisely control the generation of 3D geometry in terms of shape and appearance. In this paper, we introduce a controllable generative model (i.e. \textbf{CtrlNeRF}) that uses a single MLP network to represent multiple scenes with shared weights. Consequently, we manipulated the shape and appearance codes to realize the controllable generation of high-fidelity images with 3D consistency. Moreover, the model enables the synthesis of novel views that do not exist in the training sets via camera pose alteration and feature interpolation. Extensive experiments were conducted to demonstrate its superiority in 3D-aware image generation compared to its counterparts.
A Generative Model for Parts-based Object Segmentation
The Shape Boltzmann Machine (SBM) [1] has recently been introduced as a stateof-the-art model of foreground/background object shape. We extend the SBM to account for the foreground object's parts. Our new model, the Multinomial SBM (MSBM), can capture both local and global statistics of part shapes accurately. We combine the MSBM with an appearance model to form a fully generative model of images of objects. Parts-based object segmentations are obtained simply by performing probabilistic inference in the model. We apply the model to two challenging datasets which exhibit significant shape and appearance variability, and find that it obtains results that are comparable to the state-of-the-art. There has been significant focus in computer vision on object recognition and detection e.g.
Generalizable Imitation Learning Through Pre-Trained Representations
Chang, Wei-Di, Hogan, Francois, Meger, David, Dudek, Gregory
In this paper we leverage self-supervised vision transformer models and their emergent semantic abilities to improve the generalization abilities of imitation learning policies. We introduce BC-ViT, an imitation learning algorithm that leverages rich DINO pre-trained Visual Transformer (ViT) patch-level embeddings to obtain better generalization when learning through demonstrations. Our learner sees the world by clustering appearance features into semantic concepts, forming stable keypoints that generalize across a wide range of appearance variations and object types. We show that this representation enables generalized behaviour by evaluating imitation learning across a diverse dataset of object manipulation tasks. Our method, data and evaluation approach are made available to facilitate further study of generalization in Imitation Learners.
Asset2Vec: Turning 3D Objects into Vectors and Back
At Datagen, where I currently work as the Head of AI research, we create synthetic photorealistic images of common 3D environments, for the purpose of training computer vision algorithms. For example, if you want to teach a house robot to navigate through a messy bedroom like the one below, it will take you quite some time to collect real images for a large-enough training set [people don't usually like when outsiders enter their bedroom, and they definitely won't appreciate taking pictures of their mess]. We generated the above image using a graphic software. Once we build (in the software) the 3D models of the environment and everything in it, we use its ray-tracing renderer to generate an image of the scene from any camera view point we like, under any light conditions we want. Now if you think collecting real images is hard, wait till you try to label these images.
Deep4D: A Compact Generative Representation for Volumetric Video
This paper introduces Deep4D a compact generative representation of shape and appearancefrom captured 4D volumetric video sequences of people. 4D volumetric video achieves highlyrealistic reproduction, replay and free-viewpoint rendering of actor performance from multipleview video acquisition systems. A deep generative network is trained on 4D video sequencesof an actor performing multiple motions to learn a generative model of the dynamic shapeand appearance. We demonstrate the proposed generative model can provide a compactencoded representation capable of high-quality synthesis of 4D volumetric video with two ordersof magnitude compression. A variational encoder-decoder network is employed to learn anencoded latent space that maps from 3D skeletal pose to 4D shape and appearance. Thisenables high-quality 4D volumetric video synthesis to be driven by skeletal motion, includingskeletal motion capture data. This encoded latent space supports the representation of multiplesequences with dynamic interpolation to transition between motions. Therefore we introduceDeep4D motion graphs, a direct application of the proposed generative representation. Deep4Dmotion graphs allow real-tiome interactive character animation whilst preserving the plausiblerealism of movement and appearance from the captured volumetric video. Deep4D motion graphsimplicitly combine multiple captured motions from a unified representation for character animationfrom volumetric video, allowing novel charact...
A Generative Model for Parts-based Object Segmentation
Eslami, S., Williams, Christopher
The Shape Boltzmann Machine (SBM) has recently been introduced as a state-of-the-art model of foreground/background object shape. We extend the SBM to account for the foreground object's parts. Our model, the Multinomial SBM (MSBM), can capture both local and global statistics of part shapes accurately. We combine the MSBM with an appearance model to form a fully generative model of images of objects. Parts-based image segmentations are obtained simply by performing probabilistic inference in the model. We apply the model to two challenging datasets which exhibit significant shape and appearance variability, and find that it obtains results that are comparable to the state-of-the-art.
Generative Affine Localisation and Tracking
We present an extension to the Jojic and Frey (2001) layered sprite model which allows for layers to undergo affine transformations. This extension allows for affine object pose to be inferred whilst simultaneously learning the object shape and appearance. Learning is carried out by applying an augmented variational inference algorithm which includes a global search over a discretised transform space followed by a local optimisation. To aid correct convergence, we use bottom-up cues to restrict the space of possible affine transformations. We present results on a number of video sequences and show how the model can be extended to track an object whose appearance changes throughout the sequence.
Generative Affine Localisation and Tracking
We present an extension to the Jojic and Frey (2001) layered sprite model which allows for layers to undergo affine transformations. This extension allows for affine object pose to be inferred whilst simultaneously learning the object shape and appearance. Learning is carried out by applying an augmented variational inference algorithm which includes a global search over a discretised transform space followed by a local optimisation. To aid correct convergence, we use bottom-up cues to restrict the space of possible affine transformations. We present results on a number of video sequences and show how the model can be extended to track an object whose appearance changes throughout the sequence.
Generative Affine Localisation and Tracking
We present an extension to the Jojic and Frey (2001) layered sprite model which allows for layers to undergo affine transformations. This extension allows for affine object pose to be inferred whilst simultaneously learning theobject shape and appearance. Learning is carried out by applying an augmented variational inference algorithm which includes a global search over a discretised transform space followed by a local optimisation. Toaid correct convergence, we use bottom-up cues to restrict the space of possible affine transformations. We present results on a number of video sequences and show how the model can be extended to track an object whose appearance changes throughout the sequence.